Back

Nature Computational Science

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Nature Computational Science's content profile, based on 55 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.

1
DigiMus: a connectome-informed spiking framework for multi-region mouse neural-behavior modeling

Liu, Y.; Zhang, X.; Chen, X.; Hao, C.; Yao, W.; Zhang, J.; Sun, Y.; Zhang, T.

2026-06-11 neuroscience 10.64898/2026.06.09.731075 medRxiv
Top 0.1%
7.8%
Show abstract

Computational models are increasingly used to relate mouse brain structure, neural activity and behavior, but most models still learn from task data with limited constraints from biological circuit organization. Here we present DigiMus, a connectome-informed spiking framework for multi-region-capable mouse neural-behavior modeling. DigiMus combines leaky integrate-and-fire spiking dynamics with brain-region-specific motif regularization in a trainable sequence-modeling architecture, allowing directed three-node circuit motifs derived from 38,481 reconstructed neuronal morphologies across approximately 50 brain regions to guide recurrent coupling during learning. We evaluate DigiMus on 18 rule-based cognitive tasks spanning sensorimotor mapping and perceptual decision-making, and on three mouse neural decoding datasets involving auditory discrimination, fixed-interval licking and visual decoding. Across synthetic tasks, DigiMus showed stable performance relative to TCN, LSTM and Transformer baselines, with stronger advantages in more complex decision-making settings. In real neural datasets, single-region instantiations of DigiMus produced small, consistent and dataset-dependent improvements over a structure-free sequence baseline, while retaining motif-prior signatures in trained connectivity. Internal state analyses further linked task-dependent state dynamics to behavioral error patterns. These results suggest that connectome-derived structural priors can shape neural sequence models, and establish DigiMus as a modular, connectome-informed workflow for mouse neural-behavior modeling and hypothesis generation, rather than a complete digital reconstruction.

2
ProtGPT3: an Open-source family of Promptable and Aligned Protein Language Models

Garibbo, M.; Boxo Corominas, G.; Stocco, F.; Illanes Vicioso, R.; Middendorf, L.; Ferruz, N.

2026-06-08 bioengineering 10.64898/2026.06.04.730041 medRxiv
Top 0.1%
7.2%
Show abstract

Generative protein language models (pLMs) enable exploration of vast sequence spaces for protein design, but reliably controlling generation toward desired functional families remains challenging. While protein generation has broadly followed trends in NLP, two directions remain underexplored: alignment methods that optimize model behavior toward design objectives, and prompting-based control at inference time without fine-tuning. We introduce ProtGPT3, an open-source family of protein language models spanning 112M to 10B parameters and integrated with the Hugging Face ecosystem. The suite includes both single-sequence and multiple sequence alignment (MSA)-promptable models, enabling flexible conditioning for generation. Across model scales and protein families, we systematically compare supervised fine-tuning and few-shot prompting using homologous sequences. Analogous to how large language models (LLMs) are routinely aligned with user intent, we study post-training alignment in single-sequence models using sequence-complexity and structure-confidence metrics across the proteome. We find that alignment reduces low-complexity generations while preserving sequence diversity. Furthermore, we show that few-shot prompting is a competitive and more scalable alternative to supervised fine-tuning for controlled generation. In a low-data defluorinase case study, ProtGPT3-MSA achieved higher computational success rates than fine-tuned baselines and produced designs that were soluble and expressed following experimental validation. Finally, we explore the potential of inference-time compute in MSA models by introducing a homolog-based Feynman-Kac inference procedure for steering protein generation toward desired targets. All models are publicly available at https://huggingface.co/collections/AI4PD/protgpt3-family.

3
Connectome quality converges predictably to reveal optimal stopping points during proofreading

Martinez, H.; Matelsky, J.; Xenes, D.; Merfeld, K.; Cavanaugh, C.; Rivlin, P.; Smith, C. J.; Wester, B.

2026-07-04 neuroscience 10.64898/2026.06.30.735414 medRxiv
Top 0.1%
6.8%
Show abstract

Volumetric electron microscopy (EM) has become a critical approach to generating high-resolution reconstructions of brain tissue. As the size of EM volumes increase, use of automated image segmentation within the reconstruction pipeline has become essential, although it generates errors that need correction. The proofreading and correcting of these errors has since become the dominant cost driver in the pipeline, but precisely estimating the sufficient number of proofreading edits to enable meaningful scientific analyses of the reconstructed neuronal networks remains a challenge. We present a fast, computationally inexpensive way to estimate the progress of a connectomic proofreading effort without requiring a priori knowledge of ground truth. We show that simple global graph invariants converge predictably to asymptotic limits with increasing numbers of proofreading edits, informing a quantitative "pencils down" criterion for proofreading completeness. We illustrate our method on two datasets in different stages of proofreading progress, a zebrafish spinal cord and the hemibrain Drosophila melanogaster dataset. Our method reduces the uncertainty associated with the planning and prioritization of proofreading activities and enables data owners to accurately predict and budget the amount of proofreading necessary for their scientific questions.

4
TaxoFormer: Hierarchical Transformer for Predicting the Full Taxonomic Lineage of Protein Sequences

Parsa, M.; Azimian, K.; Wei, K. Y.

2026-06-09 synthetic biology 10.64898/2026.06.06.730618 medRxiv
Top 0.1%
5.5%
Show abstract

Predicting labels in massive, hierarchically structured output spaces is a core challenge in machine learning. In this work, we use the problem of predicting the full taxonomic lineage of a protein from its sequence as a case study for this challenge. We introduce TaxoFormer, an architecture whose primary contribution is a structured tokenization scheme that losslessly represents the entire NCBI phylogenetic tree, a graph with over 1.3 million nodes using a compact vocabulary of just 15,000 tokens. By coupling a pre-trained ESM-2 model with an autoregressive decoder and training with a standard cross-entropy objective, we test the hypothesis that a simple generative objective is sufficient to learn complex, latent structure when the output space is explicitly modeled. We show that this approach is highly effective: on a dataset of 188 million proteins, the model not only achieves accurate lineage prediction but also implicitly learns a continuous, phylogenetically-structured latent space. This work provides a scalable, alignment-free method for taxonomic annotation and demonstrates that explicitly modeling the structure of a complex output space is a powerful mechanism for learning meaningful representations.2

5
Learning shared forecast-error structure to improve ensemble forecasts of seasonal respiratory outbreaks

Qin, Y.; Du, H.; Pei, S.

2026-07-22 infectious diseases 10.64898/2026.07.20.26358508 medRxiv
Top 0.1%
5.3%
Show abstract

Real-time forecasts of seasonal respiratory outbreaks are critical for public health preparedness and healthcare planning. Multi-model ensembles, which combine predictions from individual models, have become a leading approach for operational outbreak forecasting. Their success, however, depends in part on the assumption that component models make sufficiently independent errors. Here, we examined this assumption using archived real-time forecasts for influenza hospitalizations and influenza-like illness (ILI) in the United States. We found that component models with diverse structures and calibration methods shared systematic forecast errors during epidemic growth and around epidemic peaks, reflecting the common challenge of tracking rapid changes in epidemic dynamics from real-time surveillance data. Because such shared errors cannot be fully corrected by ensembling alone, we developed a deep learning framework that learns structured residual errors from historical forecasts and uses them to correct ensemble predictions. This framework improved influenza hospitalization forecasts across horizons and geographic scales, reducing the Weighted Interval Score by up to 20\% at the national level and 12\% across states relative to official ensemble forecasts, with the largest improvements at the near-term horizon and during epidemic growth and peak periods. We further showed that learned residual structures transferred across ensembles formed from different component models, making the approach robust to changes in model participation across seasons. The framework also improved ensemble forecasts for ILI, although gains were more modest. These findings reveal a fundamental challenge in ensemble forecasting and provide a generalizable approach for improving real-time epidemic forecasts.

6
geneXplore: An Interactive Browser for X Chromosome-Wide Association Study Results

Cook, N.; Boulais-Richard, J.; Zeng, Y.; Yang, C.; Budde, J.; Taliun, D.; Gagliano Taliun, S. A.; Cruchaga, C.; Belloy, M. E.

2026-07-14 neurology 10.64898/2026.07.14.26357489 medRxiv
Top 0.1%
5.0%
Show abstract

Summary: The X chromosome comprises approximately 5% of the human genome and encodes over 800 protein-coding genes, many of which exhibit sex-differentiated expression patterns due to escape from X chromosome inactivation (XCI) mechanisms. Despite its relevance to sex differences in complex traits, the X chromosome is routinely excluded from genome-wide association studies due to analytical challenges, and when analyzed, the impact of escape from XCI or sex is limitedly explored. No dedicated, publicly accessible browser for X chromosome-wide association study (XWAS) summary statistics currently exists, creating a barrier to systematic investigation of X-linked contributions to human traits. Here, we present geneXplore, an interactive web browser based on the PheWeb2 implementation, tailored for XWAS summary statistics across 1,944 phenotypes while distinguishing random XCI (rXCI), escape from XCI (eXCI), and sex-stratified analyses. Users can explore results via interactive plots (Manhattan and Miami, PheWAS and LocusZoom), searchable tables and access to cross-database lookup, with full summary statistics available for download. Availability and Implementation: geneXplore is freely available at https://genexplore.wustl.edu/ with no registration required and will be maintained for a minimum of two years following publication. Source code is available at https://github.com/Belloy-Lab/geneXplore_XWAS_Browser under an MIT license.

7
GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler

Uzum, A. S.; Haliloglu, T.

2026-09-01 bioinformatics 10.64898/2026.08.28.747885 medRxiv
Top 0.1%
4.8%
Show abstract

Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.

8
Breaking the Synthesis Barrier for AI-Designed DNA Libraries

Sussex, S.; Borevkovic, E.; Lohmann, F.; Chen, N.; Lüthi, E.; Reddy, S. T.; Krause, A.

2026-07-07 bioengineering 10.64898/2026.07.07.736931 medRxiv
Top 0.1%
4.7%
Show abstract

Designing DNA libraries is a key challenge from drug design to protein engineering and synthetic biology. Modern generative models offer opportunities to navigate the design space and propose specific sequences predicted to be effective in-silico. Designing deterministic libraries of specific sequences is however limited by the cost of DNA synthesis -- the synthesis barrier. In contrast, high-throughput multiplexed screening can measure the function of billions of biological sequences in parallel. Harnessing this technology requires the design of randomized libraries with specific design constraints to achieve low synthesis costs. In practice, such stochastic libraries are often chosen heuristically, sacrificing control for scale. Is there a way to bridge AI-based in-silico sequence design with high-throughput experimentation? In this work, we introduce Policy Gradients for Library Design (PGLD). PGLD uses a synthesis-aware parametrization of stochastic DNA libraries and optimizes them against a specified objective function. This allows for designing massive, controlled libraries without being limited by synthesis costs. We show how PGLD enables lab-in-the-loop design of multi-round high-throughput experiments, and large-scale in-vitro DNA sampling from generative models. Finally, we use PGLD to design a library of ~10^6 unique sequences which is synthesized at a cost of ~700 USD to explore the mutation space of a broadly neutralizing influenza antibody.

9
PRIME: scalable, robust inference of mechanistic cell states from multimodal single-cell counts via probability generating functions

Li, S.; Wang, Y.; Jiang, Q.; Grima, R.; Cao, Z.

2026-06-09 systems biology 10.64898/2026.06.04.730253 medRxiv
Top 0.1%
4.7%
Show abstract

Single-cell multiomic technologies can now quantify complementary RNA species within the same cell, creating an opportunity to move beyond descriptive clustering toward mechanistically interpretable cell states. Yet most current methods depend on heuristic integration steps and become computationally burdensome at scale, limiting their ability to robustly detect subtle kinetic differences across heterogeneous populations. Here we introduce PRIME, a scalable framework for mechanistic cell-state discovery from multimodal single-cell count data. PRIME embeds multimodal measurements in a probability generating function (PGF) space, where transcriptional dynamics are encoded compactly and compared efficiently. This representation enables robust inference of latent kinetic structure and supports rapid cell grouping with a power K-means backbone that remains stable under noise, sparsity, and multimodality. Across synthetic benchmarks and experimental multimodal datasets, PRIME consistently recovers cell populations distinguished by transcriptional kinetics, outperforms conventional integration-and-clustering pipelines in robustness, and yields interpretable parameters that link observed variability to underlying regulatory mechanisms. By providing a mathematically principled yet practical route from multimodal counts to kinetic cell states, PRIME empowers biologists to uncover dynamic transcriptional regimes, dissect regulatory heterogeneity, and connect cell identity to mechanism rather than markers.

10
BraiNN: A Modern Simulator for Clinically Feasible Personalized Whole-Brain Network Modeling

Fasse, A.; Billi, C.; Garvalov, V.; Morvan, M.; Newton, T.; Kuster, N.; Neufeld, E.

2026-07-13 neuroscience 10.64898/2026.07.08.737156 medRxiv
Top 0.1%
4.4%
Show abstract

Personalized whole-brain modeling aims to transform treatment planning for neurological disorders by enabling patient-specific simulations of brain network dynamics. Neural mass models (NMMs) offer a tractable compromise between biophysical detail and computational cost and can be directly linked to macroscopic observables such as EEG. However, scaling NMMs to whole-brain networks with realistic connectivity, conduction delays, and cortical surface resolution--and fitting them to individual patient data--imposes computational demands that existing frameworks cannot meet at clinically relevant timescales. Here we introduce BraiNN, a JAX-based Python framework for large-scale neural mass modeling that achieves speedups of up to two to three orders of magnitude over existing tools by leveraging GPU/TPU-accelerated, XLA-compiled array computation. BraiNN combines a region-level Jansen-Rit network with a subject-specific cortical surface mesh of coupled neural mass models and biophysically grounded EEG forward modeling via reciprocity-based lead fields. Its fully differentiable computational graph enables a hybrid personalization pipeline that pairs Bayesian optimization for global parameter exploration with gradient-based refinement, completing EEG-driven spectral fitting of an eight-dimensional parameter space in approximately 2-3 hours on a single consumer GPU--compared to multiple days with conventional neural mass modeling software. Numerical verification against established benchmarks confirms that BraiNN faithfully reproduces canonical synchronization and bifurcation dynamics of Jansen-Rit networks. By reducing the time requirements for personalizing a high-detail whole-brain surface model from days to a few hours on consumer-grade hardware, BraiNN brings personalized brain network modeling closer to practical use in clinical contexts. We anticipate that BraiNN will serve as a foundation for patient-specific digital twins and EEG-guided neuromodulation planning.

11
cuBayes: GPU accelerated FreeBayes that achieves 1-minute whole-genome SNV calling while maintaining algorithmic semantics

Pitman, A.; Yang, C.; Qiao, Y.

2026-06-16 bioinformatics 10.64898/2026.06.12.731910 medRxiv
Top 0.1%
4.3%
Show abstract

Next-generation sequencing now produces whole-genome data in hours, but downstream variant calling remains a multi-hour to multi-day bottleneck that excludes genomic analysis from time-critical clinical settings. GPU acceleration offers a natural path forward -- variant calling is inherently parallelizable across genomic positions -- yet open-source infrastructure for porting existing algorithms to GPU hardware remains limited, leaving many widely-used tools without accelerated implementations. FreeBayes, a haplotype-based variant caller central to the 1000 Genomes Project and to multi-sample tumor evolution analyses, exemplifies this gap: it is natively single-threaded despite its algorithmic suitability for parallelization. We present cuBayes, a CUDA implementation of FreeBayes germline SNV calling that completes HG002 and HG004 2x250bp Illumina 60x whole-genome analysis in one minute (as opposed to hours if not days with manual region-based CPU parallelization) on a single NVIDIA RTX 6000 Ada GPU, while producing variant calls with 99.97% concordance to the CPU reference. cuBayes is structured around an atom/molecule architecture in which reusable functional units (BAM decompression, position-wise pileup, batch coordination) are cleanly separated from algorithm-specific logic, providing a foundation intended to support acceleration of additional sequence analysis algorithms without redundant low-level engineering.

12
Autonomous Loop Construction And Supervision For Clinician-Oriented Medical-Ai Research

Chen, X.; Jiang, X.; Shan, C.; Wang, Z.; Li, D.; Zhao, C.

2026-08-25 radiology and imaging 10.64898/2026.08.21.26361049 medRxiv
Top 0.1%
4.1%
Show abstract

Medical AI models have made a great impact on biomedical research and real-world clinical applications, but conducting interdisciplinary medical AI research remains challenging, requiring close collaboration between clinicians and AI experts. Recent advances in large language models (LLMs) and autonomous code agents present an opportunity for low cost medical AI development, where clinicians can build AI tailored to their own research questions, even without continuous support from dedicated AI experts. However, enabling code agents to autonomously tackle complex multimodal medical AI development tasks requires clinicians to construct and supervise an AI research loop with detailed technical specifics, demanding substantial expertise in AI and computer science that they often lack. To address this challenge, we introduce the Medical AI Research Loop Agent (MARLA), an agentic framework that completely abstracts the construction and supervision of medical AI research loops from clinicians. Given a clinician-defined research intent, MARLA automatically translates high-level research goals into executable hierarchical research loops, decomposes them into verifiable sub-loops, and specifies the models, datasets, tools, and evaluation protocols required for each task. During execution, MARLA coordinates specialized code agents, monitors progress, diagnoses failures, and iteratively refines research strategies based on experimental feedback to drive the research process toward optimal outcomes. We evaluate MARLA on multimodal medical AI tasks that require closed-loop conversion from high-level clinical study objectives to trained and validated AI models. The results demonstrate MARLA's ability to autonomously conduct complex medical AI research while substantially reducing the need for AI expertise.

13
A Structural Principle for Macroscopic Neural Dynamics Correlations

Wu, Q.; Wen, Q.; Liu, C.

2026-06-17 neuroscience 10.64898/2026.06.14.729168 medRxiv
Top 0.1%
4.1%
Show abstract

A central question in neuroscience is how the brains structural connectivity gives rise to its emergent, correlated dynamics. These large-scale dynamical correlations underlie functional networks that support cognitive functions. Here, we identify coupling correlation--the similarity between the input connectivity profiles of brain regions--as a key structural determinant of macroscopic neural dynamical correlation. Using dynamical mean-field theory (DMFT) and numerical simulations of random neural network models, we demonstrate that coupling correlation quantitatively governs dynamical correlation. The functional form of this structure-function mapping is dictated by the eigenvalue spectrum of the coupling correlation matrix: networks with bulk eigenspectra exhibit an exact linear relationship, whereas biologically plausible long-tailed spectra yield an approximately linear mapping except when the magnitude of coupling correlation approaches unity. Particularly, a long-tailed spectrum is necessary to reproduce the appropriate magnitude and size-invariance of coupling correlations observed in empirical data, thereby sustaining non-vanishing dynamical correlations that may support brain function in large systems. The theoretical prediction of approximate linearity is consistently validated using empirical datasets that include both structural coupling and neural dynamics in humans, mice, and Drosophila. Together, these results provide a mechanistic and quantitative framework linking macroscopic brain network structure to emergent neural dynamics--an essential step toward a theory of structure-function relationship in the brain. Significance StatementHow the brains wiring gives rise to its coordinated activity is a fundamental unsolved problem in neuroscience. Prior work has identified correlations between structural and functional connectivity, but these relationships lacked a mechanistic, first-principles explanation. Here, we derive an analytical framework using Dynamical Mean-Field Theory and random neural network models to show that a single structural statistic--coupling correlation, the similarity between the input connectivity profiles of brain regions--linearly and causally determines the magnitude of correlated neural dynamics. We further show that a long-tailed eigenvalue spectrum in biological structural connectivity is necessary to sustain the strong, size-invariant functional correlations observed across species. Validated in humans, mice, and Drosophila using multiple imaging and connectome modalities, this principle may provide a quantitative bridge between structural connectomics and emergent brain dynamics, with implications extending to a broad class of complex networked systems.

14
A Biologically Constrained Continuous-Time Framework for Long-Horizon Cognition Forecasting in Alzheimer's Disease

Deepika, P.; Sunkari, S.; Upadhyayula, S. K.; The Alzheimer's Disease Neuroimaging Initiative, ; Sundaresan, V.

2026-08-10 neurology 10.64898/2026.08.07.26359964 medRxiv
Top 0.1%
4.0%
Show abstract

Accurate long-term forecasting of cognitive trajectories across the Alzheimer's disease continuum is essential for early intervention, personalized prognosis, patient stratification, and clinical trial enrichment. Despite the promising predictive performance of recent longitudinal forecasting methods, they remain largely data-driven, struggle with irregularly sampled, incomplete longitudinal data and often neglect established disease biology, leading to biologically implausible trajectories. To address this, we propose a biologically constrained continuous-time framework for long-horizon cognition forecasting from limited baseline observations. The proposed method models the complete amyloid-tau-vascular-neurodegeneration-cognition (ATVNC) cascade using hierarchical Neural ODEs with biologically motivated monotonicity constraints. Each pathological stream is governed by a dedicated Neural ODE initialized from irregular longitudinal observations using a GRU-D encoder, capturing intrinsic disease evolution while being modulated by directed upstream pathological influences. A bounded cognition readout ensures physiologically valid cognitive score (MoCA) predictions, while teacher-student knowledge distillation improves learning from sparse longitudinal supervision. Evaluated on the ADNI dataset, the proposed framework achieves a long-horizon extrapolation MAE of 2.06 on 188 held-out participants while eliminating biologically implausible trajectory violations. It further demonstrates robust zero-shot cross-cohort generalization on OASIS-3 (MAE 2.68 on 300 participants), with fine-tuning improving MAE to 1.90. The model also supports prognostic enrichment for Alzheimer's clinical trials, achieving up to 2.70x enrichment over the cohort base rate. These results demonstrate that embedding biological disease mechanisms within continuous-time deep learning improves the accuracy, biological plausibility, and clinical utility of long-horizon cognitive forecasting. The code is publicly available at: https://github.com/PonDeepika/BEACON.

15
Towards the Virtual Amyotrophic Lateral Sclerosis Patient: Inferring Cortical Excitability through Whole-Brain Dynamical Modeling

Angiolelli, M.; Demuru, M.; Lopez, E. T.; Hashemi, M.; Ziaeemeh, A.; Rabuffo, G.; Trojsi, F.; Granata, C.; Tafuri, D.; De Luca, M.; Gallo, E.; Jirsa, V.; Depannemaecker, D.; Sorrentino, P.

2026-06-10 neurology 10.64898/2026.06.09.26354829 medRxiv
Top 0.1%
3.9%
Show abstract

Amyotrophic lateral sclerosis (ALS) is increasingly recognized as a multisystem neurodegenerative disorder in which motor-neuron degeneration is accompanied by widespread alterations in cortical dynamics. Among its most reproducible neurophysiological signatures is cortical hyperexcitability, yet how this local excitability imbalance shapes distributed whole-brain activity remains poorly understood. Here, we combined source-reconstructed resting-state MEG data, tractography-informed whole-brain modeling, and simulation-based inference to investigate whether ALS-related alterations in large-scale brain dynamics can be mechanistically explained by changes in cortical excitability. First, we characterized empirical brain dynamics using complementary features spanning regional activity amplitude and variability, functional connectivity, and avalanche-based metrics. These analyses revealed significant alterations in ALS patients relative to healthy controls, as well as associations with clinical impairment and disease staging. To mechanistically interpret these changes, we employed a reduced Wong-Wang whole-brain model in which local recurrent excitation modulates emergent large-scale neural dynamics. Simulations showed that increasing excitability systematically reproduced the empirical dynamical signatures observed in ALS. We then applied a simulation-based inference framework to estimate latent excitability parameters directly from empirical observations. Whole-brain model inversion revealed increased excitability in ALS patients compared with controls. The recovered excitability parameter was associated with disease staging, supporting its clinical relevance as a model-derived descriptor of ALS progression. Finally, by extending the model to estimate frontal and non-frontal excitability separately, we found that ALS-related alterations were predominantly associated with increased frontal excitability, whereas non-frontal regions appeared comparatively less affected. The recovered parameters related to disease staging. Together, these findings provide a mechanistic framework linking altered large-scale brain dynamics in ALS to selective cortical hyperexcitability, explaining how local excitability changes can give rise to global network reorganization. More broadly, they show how computational model inversion can recover latent multiscale pathophysiological processes from empirical neural recordings, offering a non-perturbative alternative to complex experimental paradigms typically required to causally probe local-to-global mechanisms.

16
A Scalable Framework for Harmonized mtDNA Analysis Across Diverse Biobanks

Schecter, D. R.; Lee, S. S.; Vimal, T.; Lahoti, Y.; Goncalves, V. F.; Retallick-Townsley, K.; Pang, J.; Guvenek, A.; Preuss, M.; Tinker, R. J.; Morava, E.; Kozicz, T.; Hirano, M.; Ganesh, J.; Naini, A.; Liang, J.; Davis, L.

2026-08-25 genetic and genomic medicine 10.64898/2026.08.21.26361041 medRxiv
Top 0.1%
3.9%
Show abstract

Mitochondrial DNA (mtDNA) is increasingly recognized as an important contributor to human disease and population variation, yet most genomic biobanks do not provide standardized mtDNA variant datasets despite abundant mitochondrial sequencing reads in existing whole exome and whole-genome sequencing data. We developed a scalable framework based on the Mitoverse mtDNA Server 2 Fusion workflow to generate harmonized, analysis-ready mtDNA resources across diverse biobank infrastructures. The framework was implemented in the Mount Sinai Million Health Discoveries Program (54,151 participants) using the native Nextflow workflow and adapted for the All of Us Research Program (197,361 participants) using a custom cloud implementation that preserved the same analytical strategy. Across 251,512 participants, the framework generated standardized mtDNA datasets containing 12.9 million variant observations suitable for downstream genomic and electronic health record linked analyses. This framework enables reproducible, population-scale mitochondrial genomics across institutional and national biobanks without requiring additional sequencing or development of new variant calling methods.

17
Incorporation of single-neuron projectome-based connectivity motifs enhances the cortex-specific performance of artificial neural networks

Sun, Y.; Yao, W.; Zhang, J.; Song, W.; Zhao, X.; Hao, C.; Chen, X.; Zeng, S.; Jia, S.; Yang, Y.; Chen, X.; Xiao, X.; Poo, M.-m.; Sun, Y.; Xu, B.; Zhang, T.

2026-06-17 neuroscience 10.64898/2026.06.12.732007 medRxiv
Top 0.1%
3.9%
Show abstract

The organizational principles of natural neural networks could inspire the new architecture design of artificial neural networks (ANNs). Analysis of single-neuron connectomes of mouse brains revealed distinct profiles of three-node connectivity motifs in various cortical areas and hippocampal formation. A connectome-informed neural network algorithm ("CINA") was developed to incorporate natural connectivity motifs into ANN algorithms represented by recurrent neural network (RNN) and transformer-based large language model (LLM). We found that incorporation of the average profile of cortical motifs improved the RNNs performance in noise-resistant categorization and motor learning benchmark tasks, as compared with RNNs with random connectivity. Notably, incorporating cortex-specific motifs further elevated the RNNs performance in tasks related to the cortical function, and this effect was enhanced by artificially increasing the bias in the motif profile. Similar experimental results were verified on an LLM using Motif-Transformer for natural language question answering and brain-signal decoding tasks. Graph-theoretic analyses showed that incorporating natural motifs drove the emergence of modular and small-world properties in ANNs. Together, we demonstrated not only connectome-inspired optimization of ANN architecture but also functional significance of specific motif profiles in various cortices.

18
From homeostasis to credit assignment: a signed-XOR connectomic motif for local directional error signalling

Pena Fernandez, M.; Gonzalez Rios, A.; Lloret Iglesias, L.; Marco de Lucas, J.

2026-06-09 neuroscience 10.64898/2026.06.05.730322 medRxiv
Top 0.1%
3.5%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWBiological neural circuits are widely thought to require local error signals that tell synapses not only that a prediction is wrong, but also in which direction to change. We previously proposed that a six-neuron XOR motif acts as a homeostatic comparator: matched sensory and predictive signals cancel locally, whereas mismatches propagate an error signal. We also showed that a shallow autoencoder can learn MNIST using a signed-XOR learning rule with local decoder errors and random feedback alignment, without gradient backpropagation. Here we introduce the signed-XOR motif, an eight-neuron, twelve-edge directed signed circuit that extends the XOR comparator with two feedback channels of opposite neurotransmitter identity. By construction, the motif can convert a binary mismatch into directional error signalling, with one pathway encoding potentiation and the other depression, while respecting Dales principle. We provide open-source tools to enumerate the motif at connectome scale and test its enrichment against degree- and sign-preserving null models. The motif is enriched 24.3x in C. elegans (Z = 52.2), significantly enriched in 59/80 FlyWire Drosophila neuropils including AVLP_L (13.9x, Z = 94.4), and strongly enriched in layers 2/3-5 of a biophysically detailed mouse primary visual cortex model (global 315x; per-pivot medians up to 852 x) while absent from layer 6. The same layer-specific pattern is found in the axon-proofread subset of the EM-reconstructed MICrONS connectome. A Brian2 leaky integrate-and-fire implementation reproduces the signed-XOR truth table, remains robust to Poisson drive, produces a graded signed error, and requires a fast-spiking parvalbumin-like pivot. These results identify signed-XOR as a recurrent connectomic pattern compatible with local homeostatic error cancellation and directional credit-assignment signals. Author SummaryHow does a brain decide which of its connections to adjust when it makes a mistake? Unlike an artificial network, it has no global error signal supplied from outside: each connection can react only to the neurons it directly touches. We ask whether a small, repeating wiring pattern could provide such a local correction signal. The pattern we study, the signed-XOR motif, compares an incoming signal with the brains own prediction of it. When the two agree, the circuit stays quiet, so already-expected activity is not relayed onward. When they disagree, it does more than flag an error: it also indicates the direction of the fix, routing it through two separate channels, one meaning "strengthen", the other "weaken", consistent with the biological rule that each neuron acts with a single sign. We provide open software to search for this pattern in three nervous systems, a worm, a fly, and a detailed model of mouse visual cortex, and find it more often than chance wiring predicts, with a striking layer-specific distribution in cortex. We also simulated the eight-cell circuit with realistic spiking neurons and confirmed that it can perform the computation, but only when its inhibitory cell is a fast-spiking type like those concentrated in the enriched layers. We do not claim that any brain uses this circuit to learn or memorize. What we provide is a specific motif that could deliver a local, directional error signal that may be useful for a neuromorphic implementation.

19
Unlocking Your Programmable and Creative RNA Sequence Designer with RDiffusion

Wang, J.; Dong, J.; Li, T.; Yang, L.; yin, J.; Chen, J.; Dong, Y.; Li, J.; Tan, C.

2026-06-13 bioengineering 10.64898/2026.06.13.732023 medRxiv
Top 0.1%
3.4%
Show abstract

As a cornerstone of the central dogma, RNA has both witnessed and actively shaped three billion years of evolution. Over this vast timescale, a remarkable diversity of RNA molecules has emerged, executing functions that extend far beyond traditional roles in information transfer. In the post-genomic era, while we have cataloged tens of millions of non-coding RNA sequences and functionally annotated millions, this knowledge merely scratches the surface of the vast and enigmatic RNA sequence space. Here, we introduce RDiffusion, a comprehensive generative model designed to extensively explore this RNA universe. RDiffusion is a diffusion-based framework that, conditioned on diverse biological features, such as desired function, family type, secondary structure, tertiary structure, or binding proteins--can guide the generation of novel RNA sequences tailored to specific specifications. We evaluate RDiffusion across a broad spectrum of RNA design tasks and find that it not only surpasses all baseline methods in design success rate and sequence diversity but also achieves state-of-the-art performance on downstream tasks, functioning as a powerful RNA foundation model. To translate RDiffusion into disease applications, we targeted osteoarthritis (OA) as a prime paradigm, utilizing the RDiffusion to perform de novo design of novel miRNA sequences guided by a customized, data-driven seed selection and screening pipeline. While these designed candidates are currently undergoing rigorous biological experimental validations, the finalized evaluation data will be comprehensively integrated and presented upon formal publication. By providing a unified approach to RNA design, we anticipate that RDiffusion will accelerate the programmable engineering of RNA--with profound implications for human health, drug development, and gene-editing tools, while also establishing a new standard for representation learning on RNA-related downstream tasks.

20
STR-PG: A Topology-decoupled Pangenome Framework for Scalable Short-read Genotyping of Short Tandem Repeats

YUAN, J.; XUE, Z.; TANG, H.; LIU, Y.; WANG, J.

2026-08-12 bioinformatics 10.64898/2026.08.07.743532 medRxiv
Top 0.1%
3.4%
Show abstract

Short tandem repeats (STRs) are a rich and highly polymorphic source of human genetic variation, but representing and genotyping them in pangenome graphs remains challenging. Explicitly encoding each STR allele as a separate graph path results in increasingly complex local structures as cohort diversity increases, leading to larger index sizes and requiring significant resources for graph reconstruction when new alleles are introduced. Here, we propose STR-PG, a topologically decoupled genome-wide framework that separates stable locus representation from scalable STR allele content. STR-PG uses topologically fixed pointer nodes to represent each target locus, while allele sequences, repeat counts, motif annotations, and population frequency metadata are stored in an external registry. Short reads are mapped to STR loci via syncmer-based flanking anchors, and genotyping is performed within a locus-specific candidate space using allele-level alignment likelihood and Bayesian inference. Newly supported alleles can be integrated through registry-level updates without the need to rebuild the graph structure. Evaluations using simulated whole-genome sequencing data, 1000 Genomes Project (1kGP) samples, and r real whole-exome sequencing data from matched whole-blood-cell controls demonstrate that STR-PG maintains accurate genotyping results across various STR classes, reproduces expected population structures, and substantially reduces the computational cost of integrating additional alleles. STR-PG provides a compact and scalable framework for population-scale STR analysis using short-read sequencing.